Data Mining Using Relational Database Management Systems

نویسندگان

  • Beibei Zou
  • Xuesong Ma
  • Bettina Kemme
  • Glen Newton
  • Doina Precup
چکیده

Software packages providing a whole set of data mining and machine learning algorithms are attractive because they allow experimentation with many kinds of algorithms in an easy setup. However, these packages are often based on main-memory data structures, limiting the amount of data they can handle. In this paper we use a relational database as secondary storage in order to eliminate this limitation. Unlike existing approaches, which often focus on optimizing a single algorithm to work with a database backend, we propose a general approach, which provides a database interface for several algorithms at once. We have taken a popular machine learning software package, Weka, and added a relational storage manager as back-tier to the system. The extension is transparent to the algorithms implemented in Weka, since it is hidden behind Weka’s standard main-memory data structure interface. Furthermore, some general mining tasks are transfered into the database system to speed up execution. We tested the extended system, refered to as WekaDB, and our results show that it achieves a much higher scalability than Weka, while providing the same output and maintaining good computation time.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Integrating XML Data with Relational Databases

XML technology is pushing the world into the ecommerce era. Relational database systems, today's dominant data management tool for business, must be able to accommodate the XML data, since collecting, analyzing, mining and managing that data will be tremendously important tasks. In this paper, we investigate the problem of managing XML data in relational database systems. We are speci cally con...

متن کامل

Performance Evaluation and Analysis of K-way join variants for Association Rule Mining

Data mining aims at discovering important and previously unknown patterns from the dataset in the underlying database. Database mining performs mining directly on data stored in relational database management systems (RDBMSs). The type of underlying database can vary and should not be a constraint on the mining process. Irrespective of the database in which data is stored, we should be able to ...

متن کامل

R/S Interfaces to Databases

Connectivity is an increasingly important part of statistical computing, and interfaces to databases are becoming important both in large-scale data mining applications and from the use of smaller personal databases. In this paper we present current work and future plans on interfacing the S language (R and S-PLUS) to databases, in particular to relational database management systems (DBMS).

متن کامل

Association Rule Mining of Relational Data

Most data of practical relevance are structured in more complex ways than is assumed in traditional data mining algorithms, which are based on a single table. The concept of relations allows for discussing many data structures such as trees and graphs. Relational data have much generality and are of significant importance, as demonstrated by the ubiquity of relational database management system...

متن کامل

Set-Oriented Indexes for Data Mining Queries

One of the most popular data mining methods is frequent itemset and association rule discovery. Mined patterns are usually stored in a relational database for future use. Analyzing discovered patterns requires excessive subset search querying in large amount of database tuples. Indexes available in relational database systems are not well suited for this class of queries. In this paper we study...

متن کامل

Database Queries, Data Mining, and OLAP

Modern, commercially available relational database systems now routinely include a cadre of data retrieval and analysis tools. Here we shed some light on the interrelationships between the most common tools and components included in today’s database systems: query language engines, data mining components, and online analytical processing (OLAP) tools. We do so by pairwise juxtaposition, which ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2006